Skip to main content

Data Architecture

Data architecture decides how health information is modelled, stored, moved and made usable. It is distinct from interoperability, which is about moving data between organisations, and from analytics tools, which are what you point at the result.

The recurring failure in health is a single store asked to do two incompatible jobs: serve clinical care in milliseconds and answer population questions across years. Separating those is the first architectural decision.


Operational and analytical are different systems​

OperationalAnalytical
Question"What is this patient's current medication?""How has coverage changed by district over three years?"
ScopeOne patient, nowWhole population, over time
LatencyMillisecondsMinutes to hours is fine
ModelNormalised, patient-centric — FHIR, openEHRDenormalised, dimensional or event-based
CorrectnessMust reflect the current truthMust reflect a consistent point in time
AvailabilityCare stops without itReports are late
Consequence of a bad queryA clinician waitsA dashboard is slow

Running analytics against the production clinical database is the most common cause of clinical system outages that have nothing to do with clinical load. Separate them, and be explicit about how data flows from one to the other.


The layers​

Point-of-service systems (EMR, LMIS, CHW, lab)
│ events, batches, bulk export
▼
┌────────────────────────┐
│ Operational data store │ normalised, current, patient-centric
│ (SHR / FHIR server) │
└───────────┬────────────┘
│ CDC, FHIR Bulk Data, scheduled extract
▼
┌────────────────────────┐
│ Raw / landing zone │ as received, immutable, dated
└───────────┬────────────┘
│ cleanse, conform, map terminology
▼
┌────────────────────────┐
│ Conformed layer │ shared dimensions: patient, facility,
│ │ provider, time, geography, product
└───────────┬────────────┘
│ aggregate, model
▼
┌────────────────────────┐
│ Marts / indicators │ DHIS2, BI, surveillance, research extracts
└────────────────────────┘

The raw layer is immutable. When a transformation turns out to be wrong — and one will — you reprocess from raw rather than discovering the original data is gone.


Models​

Canonical model — the single agreed representation used for exchange. In most health ecosystems this is FHIR against national profiles. Its purpose is to prevent every pair of systems negotiating its own format.

Logical model — the domain concepts and relationships, independent of storage. Patient, encounter, episode of care, observation, facility.

Physical model — tables, indexes, partitions. Different in the operational store and the warehouse, deliberately.

A note on canonical models: they work as an exchange contract and fail as a universal internal model. Forcing every system to store data in the canonical shape produces the worst of both — see the FHIR facade pattern for the alternative.

OMOP CDM deserves mention: a common data model from the OHDSI community, designed specifically for observational analysis at scale, with a mature terminology mapping layer. Where the goal is comparative research across institutions, it is more suitable than a FHIR-shaped warehouse. FHIR-to-OMOP pipelines are an established pattern.


Master data​

The data that everything else references, and that must be identical everywhere. In health these are the registries: patient, facility, provider, product, plus terminology.

Rules:

  1. One system owns each master domain. Everything else holds a copy and knows it is a copy.
  2. Copies are refreshed from a change feed, not maintained by hand.
  3. Historical values are retained, so old records remain interpretable — see facility registry on versioning.
  4. The warehouse's conformed dimensions come from the registries, not from whatever the source system happened to send.

Point 4 is what makes cross-system analysis possible. A warehouse whose facility dimension is built from the strings in each source's data cannot join them.


Storage patterns​

PatternWhat it isFits
Operational data storeCurrent-state, normalised, queryableThe shared health record
Data warehouseStructured, modelled, governedRoutine indicators, established reporting
Data lakeRaw files at low costLanding zone, imaging, unstructured, bulk exports
LakehouseTable formats over object storage with transactions and schemaThe pragmatic default for a new analytics platform
Event storeAppend-only log of what happenedSurveillance, audit, reconstruction

A data lake without a catalogue and governance becomes an expensive place to lose data. Every dataset needs an owner, a schema, a documented meaning and a retention rule before it lands, not after.


Data quality​

Health data quality is not an attribute of the database; it is a property of the process that created it. Measure it and feed the results back to the people who capture it.

DimensionHealth example
CompletenessFacilities reporting this month; fields populated per record
TimelinessDays between event and record; reporting lateness
AccuracyValues within plausible ranges; birth weight of 40 kg
ConsistencyThe same indicator agreeing between the EMR and the HMIS
UniquenessDuplicate patient rate — see MPI
ValidityCodes drawn from the bound value set; dates in the future
IntegrityEvery encounter references a facility that exists

The WHO Data Quality Review framework and the DHIS2 data quality tooling provide established, comparable definitions for routine health data — use them rather than inventing local metrics.

Publish the metrics per facility and per district. Data quality improves when it is visible and attributed, and essentially never improves as a result of a cleaning script.


Lineage and metadata​

Three questions the architecture must be able to answer:

  1. Where did this number come from, through which transformations?
  2. If this source changes, what breaks?
  3. What does this field actually mean, and who decides?

That requires a data dictionary (business definitions, owners, value sets), a data catalogue (technical inventory, schemas, ownership), and lineage (column-level where feasible, dataset-level at minimum).

The pragmatic version for a small team: transformations in version control, a dictionary maintained alongside the indicator definitions, and a rule that no dataset is published without an owner. Tooling (OpenMetadata, DataHub, Amundsen, OpenLineage) is worth adopting once the manual version is being outgrown, not before.


Health-specific modelling concepts​

ConceptWhy it needs a decision
Patient-centric vs encounter-centricDetermines whether longitudinal questions are answerable at all
Episode of careA pregnancy, a TB course — the unit that makes continuity legible; rarely modelled explicitly, and always needed
Longitudinal recordRequires stable identity across time and systems
Event time vs record timeData recorded weeks later is normal in community health; reports must be able to use either
Corrections and retractionsA result is amended, a diagnosis retracted. Overwriting loses the audit trail; the record must be versioned
Aggregation denominatorsCatchment population, target population, and where they come from — usually the most contested number in any report

In this section​


References​